Traceable Evidence-Centric Generation for Digital Forensics
LIN Guokai1, SUN Yuanyuan1, GUO Hong2, Paerhati TULAJAING1, YANG Liang1, LIN Hongfei1
1. School of Computer Science and Technology, Dalian University of Technology, Dalian 116024; 2. Shanghai Key Laboratory of Forensic Medicine, Shanghai Forensic Service Platform, Academy of Forensic Science, Shang- hai 200063
Abstract:To address the problems of vulnerable evidence boundaries, difficulty in preserving source and location information, and insufficient verifiable support for generated results in multi-source heterogeneous electronic evidence scenarios, a traceable evidence-centric generation method for digital forensics(TEC-Gen) is proposed. First, multi-source forensic data, including chat records, file records, call records and system logs, is parsed in a unified manner. Original records are regarded as the basic units to construct evidence tuples. Within each evidence tuple, the original record text and metadata, including source files, location information, timestamps and key entities, are jointly modeled to preserve the original evidence boundary and its location anchor. Then, candidate evidence is organized according to a case analysis query. Relevant records are retrieved while structural attributes, including source files, location information and key entities, are preserved, providing a unified evidence basis for subsequent generation and verification. In the generation stage, the candidate evidence and citation rules are jointly incorporated into the prompt template, and a large language model is required to attach source markers after key factual statements. Finally, the output citations are parsed and matched with valid original positions in the candidate evidence. Explicit back-links between generated statements and original evidence are established. Experimental results show that TEC-Gen outperforms closed-book generation method and standard retrieval-augmented generation method on different large language model backbones. Stable advantages are observed in text alignment, key fact coverage, and result traceability. Methodological support is provided for case summary generation, evidence localization, and assisted analysis in digital forensics.
[1] 韩丹东,姜珊.电子证据须保障合法真实关联性[EB/OL]. [2026-05-17]. http://legal.people.com.cn/n1/2019/0412/c42510-31026319.html. (Han D D, Jiang S. Electronic evidence must ensure legality, authenticity and relevance[EB/OL]. [2026-05-17]. http://legal.people.com.cn/n1/2019/0412/c42510-31026319.html [2] 陈云. 基于大模型的电子数据取证分析应用研究[J].中国安防, 2025(12): 32-37. (Chen Y.Research on the application of big model-based digital forensics analysis[J]. China Security and Protection, 2025(12): 32-37.) [3] Mahar M A, Raza A, Uddin Shaikh Z, et al. Transformative role of LLMs in digital forensic investigation: exploring tools, challenges, and emerging opportunities[J]. VAWKUM Transactions on Compu-ter Sciences, 2025, 13(1): 217-229. [4] Yin Z P, Wang Z C, Xu W F, et al. Digital forensics in the age of large language models[EB/OL].[2026-05-17].https://arxiv.org/pdf/2504.02963. [5] Sharma B, Ghawaly J,McCleary K, et al. ForensicLLM: a local large language model for digital forensics[J/OL]. Forensic Science International(Digital Investigation), 2025, 52. https://doi.org/10.1016/j.fsidi.2025.301872. [6] Michelet G, Henseler H, van Beek H, et al. Fine-tuning large language models for digital forensics: case study and general recommendations[J]. Digital Threats: Research and Practice, 2025, 6(4): 1-18. [7] Selamat S R, Ahmad S S S, Masud M Z, et al. TRACEMAP: a traceability model for the digital forensics investigation process[C]//Proceedings of the IEEE Conference on Application, Information and Network Security. Washington, USA: IEEE, 2017: 25-30. [8] Lewis P, Perez E, Piktus A, et al.Retrieval-augmented generation for knowledge-intensive NLP tasks[C]//Proceedings of the 34th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2020: 9459-9474. [9] 王鑫林,李岩,马超凡,等.检索增强生成(RAG)综述:方法与应用[EB/OL]. [2026-05-17]. https://link.cnki.net/urlid/50.1075.tp.20260317.1842.019. (Wang X L, Li Y, Ma C F, ,et al. Retrieval-augmented generation(RAG): a survey of methods and applications[EB/OL]. [2026-05-17]. https://link.cnki.net/urlid/50.1075.tp.20260317.1842.019.) [10] 王素,杜志淳.生成式人工智能赋能电子数据鉴定:表征、挑战与进路[J].中国司法鉴定, 2025(3): 83-91. (Wang S, Du Z C.Generative artificial intelligence empowers di-gital forensics: representation, challenges and approaches[J]. Chinese Journal of Forensic Sciences, 2025(3): 83-91.) [11] 李冰,崔航源,易丽瑾.生成式人工智能等新一代人工智能在司法鉴定领域的应用发展初探[J].中国司法鉴定, 2026(2): 53-61. (Li B, Cui H Y, Yi L J.Preliminary study on the application and development of new-generation artificial intelligence including ge-nerative artificial intelligence in forensic appraisal[J]. Chinese Journal of Forensic Sciences, 2026(2): 53-61.) [12] Wallat J, Heuss M, de Rijke M, et al. Correctness is not faithfulness in retrieval augmented generation attributions[C]//Procee-dings of the International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval. New York, USA: ACM, 2025: 22-32. [13] Scanlon M, Breitinger F, Hargreaves C,et al. ChatGPT for digital forensic investigation: the good, the bad, and the unknown[J/OL]. Forensic Science International(Digital Investigation), 2023, 46. https://doi.org/10.1016/j.fsidi.2023.301609. [14] Chikul P, Bahsi H, Maennel O.An ontology engineering case study for advanced digital forensic analysis[C]//Proceedings of the 10th International Conference on Model and Data Engineering. Berlin, Germany: Springer, 2021: 67-74. [15] 袁乐,刘绍华,王禹,等.大语言模型检索增强生成优化技术研究综述[J].计算机学报, 2026, 49(2):383-422. (Yuan L, Liu S H, Wang Y, et al. A comprehensive review of optimization techniques in retrieval-augmented generation for large language models[J]. Chinese Journal of Computers, 2026, 49(2): 383-422.) [16] 罗文培,黄德根.大模型增强的跨模态图文检索方法[J].小型微型计算机系统, 2025, 46(7): 1544-1553. (Luo W P, Huang D G.Enhanced cross-modal image-text retrieval with large models. Journal of Chinese Computer Systems, 2025, 46(7): 1544-1553.) [17] Manakul P, Liusie A, Gales M.SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models[C]//Proceedings of the Conference on Empirical Methods in Na-tural Language Processing. Stroudsburg, USA: ACL, 2023: 9004-9017. [18] Gao L Y, Dai Z Y, Pasupat P, et al. RARR: researching and revising what language models say, using language models[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics(Long Papers). Stroudsburg, USA: ACL, 2023: 16477-16508. [19] Michelet G, Breitinger F. ChatGPT, Llama, can you write my report an experiment on assisted digital forensics reports written using (local) large language models[J/OL]. Forensic Science International(Digital Investigation), 2024, 48. https://doi.org/10.1016/j.fsidi.2023.301683. [20] Villegas-Ch W, Gutierrez R, Navarro A M. Artificial intelligence techniques for enhancing accuracy and efficiency in digital forensic analysis[J/OL]. Discover Artificial Intelligence, 2026, 6(1). https://doi.org/10.1007/s44163-025-00729-4. [21] 林哲旭,陈汉林,刘漳辉,等.基于Hyperledger Fabric的数据可信共享平台[J].小型微型计算机系统, 2025, 46(1): 189-199. (Lin Z X, Chen H L, Liu Z H, et al. Data trusted sharing platform based on Hyperledger Fabric. Journal of Chinese Computer Systems, 2025, 46(1): 189-199.) [22] Aizawa A.An information-theoretic perspective of TF-IDF mea-sures[J]. Information Processing and Management, 2003, 39(1): 45-65. [23] Roberts A, Raffel C, Shazeer N.How much knowledge can you pack into the parameters of a language model[C]//Proceedings of the Conference on Empirical Methods in Natural Language Proce-ssing. Stroudsburg, USA: ACL, 2020: 5418-5426. [24] Reuter M, Lingenberg T, Liepina R, et al. Towards reliable retrieval in RAG systems for large legal datasets[C]//Proceedings of the Natural Legal Language Processing Workshop. Stroudsburg, USA: ACL, 2025: 17-30. [25] DeepSeek-AI. DeepSeek-V3.2: pushing the frontier of open large language models[EB/OL]. [2026-05-17]. https://arxiv.org/pdf/2512.02556. [26] GLM-5 team. GLM-5: from vibe coding to agentic engineering[EB/OL]. [2026-05-17]. https://arxiv.org/pdf/2602.15763. [27] Qwen team. Qwen3 technical report[EB/OL]. [2026-05-17]. https://arxiv.org/pdf/2505.09388.